Papers with linguistic processing
lingvis.io - A Linguistic Visual Analytics Framework (P19-3)
Copied to clipboard
Mennatallah El-Assady, Wolfgang Jentner, Fabian Sperrle, Rita Sevastjanova, Annette Hautli-Janisz, Miriam Butt, Daniel Keim
| Challenge: | Using a modular framework, linguistic visual analytics applications can be rapidly prototypized using a web-based framework. |
| Approach: | They propose a modular framework for rapid prototyping of linguistic, web-based, visual analytics applications. |
| Outcome: | The proposed framework supports rapid prototyping of linguistic, web-based, visual analytics applications. |
Visio-Linguistic Brain Encoding (2022.coling-1)
Copied to clipboard
| Challenge: | Existing studies have failed to explore co-attentive multi-modal modeling for visual and text reasoning. |
| Approach: | They propose to use image and multi-modal Transformers to reconstruct fMRI brain activity . they use two popular datasets to study visual and text reasoning . |
| Outcome: | The proposed model outperforms existing models on two popular datasets . the results raise the question whether visual processing is affected implicitly by linguistic processing . |
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation (2026.acl-demo)
Copied to clipboard
| Challenge: | Existing document translation pipelines face a tension between linguistic processing and layout preservation. |
| Approach: | They propose a framework for layout-preserving PDF translation that decouples visual layout metadata from semantic content. |
| Outcome: | The proposed framework improves layout fidelity, visual aesthetics, and terminology consistency over representative baselines while maintaining competitive translation precision. |
RusConText Benchmark: A Russian Language Evaluation Benchmark for Understanding Context (2025.acl-srw)
Copied to clipboard
| Challenge: | a new context understanding benchmark is proposed for short-context understanding in Russian . the benchmarks focus on broad reasoning tasks or long-concept comprehension, but are limited in their ability to perceive subtle nuances of context. |
| Approach: | They propose a new benchmark for evaluating short-context understanding in Russian . they propose to use four tasks to assess model performance from a specific perspective . |
| Outcome: | The proposed benchmark is adapted to Russian-language data. |
Evaluating Language Tools for Fifteen EU-official Under-resourced Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Evaluation of language tools available for 15 EU-official under-resourced languages . evaluation of NERC systems was problematic because of lack of universally or cross-lingually applicable named entities classification scheme. |
| Approach: | They evaluate language tools available for 15 EU-official under-resourced languages . they focus on existing NLP platforms that provide models for under-represented languages - stanton core, nl cube, uDPipe . |
| Outcome: | The evaluation of language tools for 15 under-resourced languages is reproducible . the results are below what was reported in the literature and in some cases even better than the ones reported previously. |
Interpretability of Language Models via Task Spaces (2024.acl-long)
Copied to clipboard
| Challenge: | linguistic interpretability is a method used to assess language models' ability to interpret outputs. |
| Approach: | They propose a method to assess LMs' language conceptualisations by 'similarity probing' and a technique to fine tune them via gradient differentials to disentangle the learning signals of linguistic phenomena. |
| Outcome: | The proposed method generalises larger models to overarching general concepts for linguistic tasks, and the generalisation patterns are stable throughout training and not marked by incisive stages. |
FLAG-TRADER: Fusion LLM-Agent with Gradient-based Reinforcement Learning for Financial Trading (2025.findings-acl)
Copied to clipboard
Guojun Xiong, Zhiyang Deng, Keyi Wang, Yupeng Cao, Haohang Li, Yangyang Yu, Xueqing Peng, Mingquan Lin, Kaleb E Smith, Xiao-Yang Liu, Jimin Huang, Sophia Ananiadou, Qianqian Xie
| Challenge: | Large language models (LLMs) have impressive reasoning capabilities in financial tasks, but struggle with multi-step, goal-oriented scenarios in interactive financial markets. |
| Approach: | They propose a framework that integrates large language models with gradient-driven reinforcement learning (RL) policy optimization. |
| Outcome: | The proposed framework improves performance in trading and other financial domain tasks. |
Konidioms Corpus: A Dataset of Idioms in Konkani Language (2024.lrec-main)
Copied to clipboard
| Challenge: | Konkani is a low-resource language spoken by 2.5 million speakers . idiomatic sense processing is challenging due to the nature of idioms . |
| Approach: | They propose to use crowdsourced idiomatic sentence identification to build a corpus for idioms in the Konkani language. |
| Outcome: | The proposed corpus consists of 6520 sentences written in the Konkani language. |